Papers with explanation methods
Explanation in the Era of Large Language Models (2024.naacl-tutorials)
Copied to clipboard
| Challenge: | Explanation has long been a part of communication, where humans use language to elucidate each other and transmit information about mechanisms of events. |
| Approach: | They review the opportunities and challenges of explanations in the era of large language models and examine how they can be used to generate explanations. |
| Outcome: | The proposed methods are based on the models of large language models (LLMs) and their opaque nature. |
Interpreting Language Models with Contrastive Explanations (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing explanation methods conflate evidence for various features to predict a token . existing explanation methods are less interpretable for human understanding . |
| Approach: | They propose to explain language models contrastively by looking for salient input tokens that explain why the model predicted one token instead of another. |
| Outcome: | The proposed explanations are better than non-contrastive explanations for language models . they show that contrastive explanations improve simulability for human observers . |
Considering Likelihood in NLP Classification Explanations with Occlusion and Language Modeling (2020.acl-srw)
Copied to clipboard
| Challenge: | Existing explanation methods produce invalid or syntactically incorrect data, neglecting the improved abilities of recent NLP models. |
| Approach: | They propose an explanation method that combines occlusion and language models to sample valid and syntactically correct replacements with high likelihood, given the context of the original input. |
| Outcome: | The proposed method can sample valid and syntactically correct replacements with high likelihood, given the context of the original input. |
Evaluating neural network explanation methods using hybrid documents and morphosyntactic agreement (P18-1)
Copied to clipboard
| Challenge: | a number of post hoc explanation methods for deep neural networks have been proposed . due to the complexity of the DNNs they explain, these methods are necessarily approximations and come with their own sources of error. |
| Approach: | They propose two evaluation paradigms that cover two important classes of NLP problems . they propose LIMSSE, LRP and DeepLIFT as the most effective explanation methods . |
| Outcome: | The proposed methods are most effective for explaining deep neural networks in NLP . the proposed methods can explain complex models without manual annotation . |
Evaluating Explanation Methods for Neural Machine Translation (2020.acl-main)
Copied to clipboard
| Challenge: | Neural machine translation (NMT) has seen great success during recent years. |
| Approach: | They propose a metric that measures the fidelity of explanation methods on translation tasks . they use an efficient approximation to evaluate several explanation methods . |
| Outcome: | The proposed metric is efficient and can be used on translation tasks. |
PromptExplainer: Explaining Language Models through Prompt-based Learning (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing explanation methods rely on linear approximations, accentuating irrelevant input tokens. |
| Approach: | They propose a method that aligns the explanation process with the masked language modeling task of pretrained language models and leverages prompt-based learning to generate class-dependent explanations. |
| Outcome: | Extensive experiments show that PromptExplainer outperforms state-of-the-art explanation methods. |
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing explanation methods for image classification struggle to provide faithful and plausible explanations for predictions. |
| Approach: | They propose a natural language explanation method that can be applied to any CNN-based classifier without altering its training process or affecting predictive performance. |
| Outcome: | The proposed method can be applied to any CNN-based classifier without altering its training process or affecting predictive performance. |
Lifelong Explainer for Lifelong Learners (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing explanation methods are inefficient when explaining a static black-box model. |
| Approach: | They propose a Lifelong Explanation approach that continuously trains a student explainer under the supervision of a teacher on different tasks undertaken in LL. |
| Outcome: | The proposed approach can be extended to include a teacher and maintain the same level of faithfulness to the black-box model as the student explainer while being up to 102 times faster at test time. |
Alignment Rationale for Natural Language Inference (2021.acl-long)
Copied to clipboard
| Challenge: | Existing explanation methods pick prominent features, but alignments between words or phrases are more enlightening clues to explain the model. |
| Approach: | They propose a method to generate alignment rationale explanations for co-attention based models in NLI by feature selection. |
| Outcome: | The proposed method is more faithful and human-readable compared with existing methods. |
Sequential Integrated Gradients: a simple but effective method for explaining language models (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing explanation methods such as Integrated Gradients (IG) produce a path for each word of a sentence simultaneously, which can lead to sentences with no clear meaning or a significantly different meaning compared to the original one. |
| Approach: | They propose to use a sequenced integrated gradient method to fix every word in a sentence and move it along a straight path to the word of interest. |
| Outcome: | The proposed method improves on Integrated Gradients (IG) and DIG, but can produce sentences with different meanings than the original one. |
Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior? (2020.acl-main)
Copied to clipboard
| Challenge: | a new study examines the impact of algorithmic explanations on simulatability of machine learning models . a model is simulatable when a person can predict its behavior on new inputs . |
| Approach: | They conduct human subject tests to isolate effect of algorithmic explanations on simulatability . they find ratings of explanations are not predictive of how helpful they are . |
| Outcome: | The results provide the first reliable estimates of how explanations influence simulatability . they show that ratings are not predictive of how helpful explanations are . |
Explainable Text Classification with LLMs: Enhancing Performance through Dialectical Prompting and Explanation-Guided Training (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing explanation methods that generate keywords may be less effective due to missing critical contextual information. |
| Approach: | They propose a new method to generate explanations for possible labels using LLMs and a dialectical prompt. |
| Outcome: | The proposed method significantly improves accuracy and explanation quality over state-of-the-art methods on multiple datasets from diverse domains. |
Quantifying Uncertainty in Natural Language Explanations of Large Language Models for Question Answering (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown strong capabilities, enabling concise, context-aware answers in question answering tasks. |
| Approach: | They propose a framework that provides valid uncertainty guarantees for LLMs . they also propose 'model-agnostic' uncertainty estimation method that maintains valid guarantees even under noise. |
| Outcome: | The proposed method provides valid uncertainty guarantees even under noise. |
Quantifying and Understanding Uncertainty in Large Reasoning Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for estimating generation uncertainty do not provide finite-sample guarantees for reasoning-answer generation. |
| Approach: | They propose a method that provides the uncertainty of the reasoning-answer structure with statistical guarantees. |
| Outcome: | The proposed method disentangles reasoning quality from answer correctness while establishing theoretical guarantees for efficient explanation methods. |
From Nodes to Narratives: Explaining Graph Neural Networks with LLMs and Graph Context (2026.acl-long)
Copied to clipboard
| Challenge: | Existing explanation methods for graph neural networks struggle to generate interpretable, fine-grained rationales. |
| Approach: | They propose a lightweight framework that uses large language models to generate interpretable explanations for GNNs. |
| Outcome: | The proposed framework generates interpretable explanations for GNN predictions using large language models. |